Back

JCO Clinical Cancer Informatics

American Society of Clinical Oncology (ASCO)

Preprints posted in the last 90 days, ranked by how well they match JCO Clinical Cancer Informatics's content profile, based on 22 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
LLM-Driven Extraction of NI-RADS and Imaging Tumor Characteristics to Enhance Oropharyngeal Cancer Survivorship Surveillance

Song, W.; Shbita, L.; Jang, I. J. H.; Starostina, O.; Lewis, R.; Sahli, A.; Floyd, W. R.; Mahin, M.; Rinsurongkawong, W.; Barbon, C. E. A.; Lai, S. Y.; Lee, J. J.; Shah, K.; Chen, M. M.; Hutcheson, K. A.; Fuller, C. D.; Moreno, A. C.

2026-06-17 radiology and imaging 10.64898/2026.06.11.26355483 medRxiv
Top 0.1%
38.2%
Show abstract

Abstract Purpose Radiologic surveillance is essential for oropharyngeal cancer (OPC) survivors, guiding recurrence detection and follow-up strategies. The Neck Imaging Reporting and Data System provides a standardized framework for post-treatment risk reporting at both the primary tumor site (pNI-RADs) and cervical lymph nodes (nNI-RADS). Comprehensive surveillance additionally requires assessment of disease status, including the primary tumor, nodal involvement, and distant metastases. These clinical results are often embedded as unstructured data within free-text radiology reports. We hypothesized that a large language model (LLM) can reliably extract NI-RADS score criteria and summarize key imaging features from unstructured radiology text, achieving high concordance with expert review. Methods Previously untreated OPC patients who received definitive cancer therapy were identified. Eligible imaging reports included post-treatment head and neck CT, MRI, or FDG PET/CT scans containing narrative and impression text. Examinations lacking narrative or impression text, containing pre-existing NI-RADS annotations, or involving non-surveillance imaging modalities were excluded. A total of 200 reports were randomly selected from 7,076 eligible examinations for manual abstraction using a three-reviewer consensus framework to establish a reference dataset. Using the Palantir Foundry Pipeline Builder, a GPT-5-based LLM was deployed to extract pNI-RADS and nNI-RADS scores, and key imaging features of disease status from these reports. Performance was evaluated using exact agreement and F1-based metrics. Results Agreement for no evidence of disease (score of 1) was 93.3% (126/135; F1 = 0.94) and 90.3% (130/144; F1 = 0.93) for pNI-RADS and nNI-RADS, respectively. For NI-RADS [≥]2, exact category agreement was 73.1% (38/52; macro-F1 = 0.75) for pNI-RADS and 64.3% (27/42; macro-F1 = 0.56) for nNI-RADS. Quadratic weighted {kappa} was 0.81 and 0.59, respectively. For post-treatment disease surveillance variables, agreement was 94.9% (149/157; F1 = 0.87) for primary tumor presence, 89.1% (164/184; F1 = 0.87) for nodal disease presence, and 94.7% (126/133; F1 = 0.70) for distant metastasis detection. Specificity was high across disease-status variables (0.95-0.99), with negative predictive values of 0.95 for primary tumor, 0.87 for nodal disease, and 0.99 for distant metastasis. Conclusions Our LLM-based information retrieval and classification approach for radiographic treatment response from unstructured, multidimensional imaging reports achieved high performance for disease exclusion and moderate performance for detecting suspected residual and/or new disease. This pipeline supports scalable and standardized surveillance data capture for longitudinal monitoring, clinical analytics, and survivorship research in head and neck oncology.

2
OncoGenRAG: Evidence-Grounded Retrieval and BioBERT Classification for Precision Oncology Variant Interpretation

Arif, A.; Filho, J. V. d. S.

2026-08-22 bioinformatics 10.64898/2026.08.20.746121 medRxiv
Top 0.1%
22.5%
Show abstract

The increasing use of tumor sequencing has intensified the need for fast, traceable interpretation of genomic variants. General-purpose large language models can produce fluent answers, but unsupported statements, weak provenance, and stale knowledge limit their suitability for clinical genomics. We developed OncoGenRAG, a research framework that combines a parameter-efficiently fine-tuned BioBERT classifier with an entity-aware retrieval system over a curated, multi-source oncology knowledge base. The reported knowledge base contains 933 harmonized records derived from CIViC, ClinVar/dbSNP, Open Targets, UniProtKB/Swiss-Prot, Ensembl Variation, and linked PubMed literature. The classifier assigns one of five labels: Pathogenic, Likely Pathogenic, Variant of Uncertain Significance, Benign, or Oncogenic; the retrieval component ranks evidence records using subword TF-IDF similarity and explicit gene, variant, and cancer-type matches. A rejection rule suppresses answers when retrieval support is below a prespecified threshold. In the authors held-out evaluation, the classifier achieved 92.40% accuracy, 93.15% weighted precision, 92.40% weighted recall, and 92.65% weighted F1 score. In a separate benchmark of 100 clinical-style queries, OncoGenRAG achieved reported Precision@1 of 94.5%, Precision@3 of 96.8%, and 100% database grounding. No hallucinated answer was observed under the study operational definition, compared with a 41.0% no-hallucination rate for the ungrounded baseline. These results should be interpreted as internal validation rather than proof of universal safety because query construction, annotator agreement, class-specific performance, calibration, and external validation data were not available for independent analysis. OncoGenRAG provides a transparent design for evidence retrieval and abstention, but it is a research prototype and must not be used to select treatment without expert review.

3
In Silico Trial Simulation with Artificial Intelligence-Generated Synthetic Control Cohorts Reproduces Results of a Randomized Controlled Trial in Acute Myeloid Leukemia

Kumar Reddy, K.; Hahn, W.; Winter, S.; Roellig, C.; Mueller-Tidow, C.; Serve, H.; Baldus, C. D.; Fransecky, L.; Schliemann, C.; Burchert, A.; Schaefer-Eckart, K.; Kaufmann, M.; Schetelig, J.; Bornhaeuser, M.; Middeke, J. M.; Eckardt, J.-N.

2026-07-16 health informatics 10.64898/2026.07.15.26358123 medRxiv
Top 0.1%
18.5%
Show abstract

Rising costs, slow accrual and molecular substratification of cancers necessitate novel clinical trial designs. We demonstrate that artificial intelligence-generated synthetic patients can replace real controls to reproduce results of the SORAML trial. Using external multimodal data from 1,377 acute myeloid leukemia (AML) patients from previous trials and a real-world registry, we fine-tuned a tabular foundation model to generate synthetic patients, reproducing clinical and genetic features and outcome associations. Synthetic patients were then matched to the original SORAML intervention group using Cox risk scores, replacing the original control and reproducing the original trial result with near-identical median event-free survival (EFS) and treatment effect (original hazard ratio [HR] 0.64, 95%-confidence interval [CI] 0.47-0.87, p=0.004; with synthetic control HR 0.66, 95%-CI 0.48-0.90, p=0.009). Our findings demonstrate that AI-generated synthetic patients can serve as statistically rigorous controls supporting novel trial designs.

4
Demographic Calibration Gaps in Breast Cancer Risk Prediction: Introducing the Demographic Calibration Gap Score

Eniolade, M.

2026-06-22 health informatics 10.64898/2026.06.17.26355900 medRxiv
Top 0.1%
18.4%
Show abstract

ABSTRACT: Most breast cancer prediction studies skip calibration reporting entirely. Fewer still examine calibration by demographic subgroup. Predicted probabilities that are systematically off for specific racial or gender groups produce biased clinical decisions, and aggregate statistics will not catch that. Objective: To introduce the Demographic Calibration Gap Score (DCGS), a metric that measures how much calibration error varies across demographic subgroups, and to show how it performs across five classifiers, four calibration conditions, and two datasets. Methods: Five classifiers were trained on the Wisconsin Diagnostic Breast Cancer dataset (n=569) and evaluated on a breast cancer cohort from MIMIC-IV (n=1,316). Three global calibration methods were applied: no calibration, Platt scaling, and isotonic regression. A fourth condition, subgroup-targeted Platt scaling, was applied to the MIMIC cohort. DCGS was computed as across racial and gender subgroups, with 95% bootstrap confidence intervals. Conformal prediction coverage and Demographic Coverage Gap (DCG) were reported. Results: On Wisconsin, all five models achieved AUROC above 0.98 and ECE below 0.12. Performance fell sharply on the MIMIC external cohort: AUROC dropped to 0.45-0.57 for base and globally calibrated variants, confirming distributional shift. DCGS exceeded the 0.05 clinical significance threshold in 28 of 40 model-calibration combinations on the race axis. Neither global Platt nor isotonic calibration reliably reduced DCGS below that threshold. Conformal coverage collapsed to roughly 25% on MIMIC, and racial DCG exceeded 0.15 for all 20 model-variant combinations. Conclusions: Reducing population-level ECE through global recalibration does not reliably close demographic calibration gaps. DCGS gives researchers a direct, standardized way to detect and report those disparities. Code and the DCGS computation library are released as open-source Python under the MIT License.

5
A ReAct Agentic AI System for Natural Language Querying and Statistical Analysis of The Cancer Genome Atlas Clinical Data

Korutla, R.; Amal, S.

2026-07-17 health informatics 10.64898/2026.07.15.26358188 medRxiv
Top 0.1%
18.2%
Show abstract

The Cancer Genome Atlas (TCGA) holds clinical data for over 11,000 patients across 33 cancer types, but access is hard because of complex file structures, heterogeneous formats, and the need for programming. We present an agentic system for natural language querying and statistical analysis of TCGA clinical data. The system uses a large language model as an autonomous ReAct agent that selects from eight computational tools, including data extraction, descriptive statistics, Kaplan-Meier survival analysis with log-rank tests, hypothesis testing, and verification against the curated TCGA Pan-Cancer Clinical Data Resource (CDR). The agent reasons about intermediate results, adapts its approach, and returns clinically contextualized responses with source attribution and auditable traces. We introduce TCGA-Agent-Bench, 440 queries across five difficulty tiers with ground truth from the independently curated TCGA-CDR, evaluated with dual metrics of numerical accuracy and clinical completeness. The system achieves 93.4% overall accuracy (100% single-patient lookups, 99.1% cohort statistics, 92.8% comparative analyses), outperforming a fixed rule-based pipeline (87.1%), a single-pass LLM (81.8%), and retrieval-augmented generation (66.9% on a subset). Most of the benchmark is answerable from the CDR alone, so we locate the extraction layer's value in fields the CDR lacks (drug treatments, TNM components, biomarkers, biospecimen metadata): on 26 queries targeting these, the full system answers 100% versus 3.8% for CDR-only. Ablations show the reasoning loop is most impactful (+9.1% accuracy, +22.0 completeness points). A tool-based agentic architecture enables accurate, auditable analysis of clinical repositories, with value driven by tool design and recovered fields rather than model scale.

6
Natural Language Processing Based Solution for Labeling Brain Metastasis Identified in Radiology Reports

Liu, T.; Han, Y. T.; Zuo, H.; Das, S.; Lin, H.-M.; Colak, E.; Istasy, M.; Ladak, A. M.; Bigenimana, J. C.; Gondara, L.; Simkin, J.; Lee, J.; Roozbeh, D.; Nichol, A. M.; Easaw, J.; Walker, E.; Yip, S.; Mou, L.; Yuan, Y.

2026-06-15 epidemiology 10.64898/2026.06.10.26355415 medRxiv
Top 0.1%
15.5%
Show abstract

Abstract Purpose: Brain metastases (BM) far exceed primary CNS tumours and constitute the majority workload for neuro-oncology care providers. Currently, the cancer registries only capture synchronous BMs, which is only a small proportion of all BMs. We aim to develop and validate a natural language processing (NLP) algorithm that identifies brain metastases in radiology reports, enabling scalable surveillance of asynchronous BMs. Methods: Using population-based cancer registry data in Alberta, Canada, we identified a cancer cohort diagnosed between 2012--2019 with follow-up to 2022. All brain/head radiology reports at and post-cancer diagnosis were identified. Reports were sampled through a multi-phase approach and manually labeled for BM presence. We trained two Bio_ClinicalBERT models on the "Findings" and "Impressions" sections, respectively, and took the maximum predicted probability as the report-level prediction. Internal and external validation used reports from the Canadian provinces of Alberta, Ontario, and British Columbia. Results: The models were trained on 1,879 samples. For internal validation, 1,833 reports from 357 patients were tested. At a probability threshold of 0.4, the model achieved a sensitivity of 0.888 and precision of 0.499. The ensemble substantially outperformed single-section models, which achieved sensitivities of only 67.8% (Findings) and 74.2% (Impressions). On external validation, sensitivity was 0.918 in Ontario and 0.726 in British Columbia, demonstrating robustness across diverse data distributions. Conclusions: An NLP-based pipeline processing both Findings and Impressions sections has been developed and validated in three Canadian provinces. It meets cancer registry operational requirements and to be implemented into the surveillance workflow in Alberta and British Columbia, providing a foundation for population-level BM surveillance.

7
Retrieval-Augmented Large Language Models for Clinically Aligned Adverse Event Coding in Acute Myeloid Leukemia Clinical Trials

Dashti, N.; Schneider, M. M. K.; Eckardt, J. N.; Fiebig, F.; Schweigler, D.; Buttner, S.; Middeke, J. M.; Bornhauser, M.; Rollig, C.; Kather, J. N.; Wiest, I. C.

2026-08-18 health informatics 10.64898/2026.08.17.26360282 medRxiv
Top 0.1%
13.2%
Show abstract

Background: Adverse event (AE) coding is essential for safety monitoring in oncology clinical trials, particularly in acute myeloid leukemia (AML), where intensive therapies are associated with frequent and heterogeneous toxicities requiring standardized MedDRA (Medical Dictionary for Regulatory Activities) coding. However, manual Low-Level Term (LLT) assignment remains labor-intensive, subjective, and difficult to scale. Although large language models (LLMs) have emerged as promising decision-support tools for automated coding, unguided zero-shot generation remains insufficient for reliable fine-grained MedDRA coding. Objective: To develop and evaluate a retrieval-augmented reasoning pipeline for clinically aligned LLT-level MedDRA coding of free-text adverse events from prospective AML clinical trials. Methods: We implemented a retrieval-augmented reasoning pipeline inspired by the retrieval-augmented generation (RAG) paradigm using LLaMA-3.3-70B-Instruct as the primary backbone and benchmarked the framework across multiple open instruction-tuned LLMs. Dense semantic retrieval first generated a constrained top-100 LLT candidate set for each AE, followed by structured LLM reasoning to select a single best-matching LLT and deterministic mapping to Preferred Term (PT) and System Organ Class (SOC) levels. The pipeline was evaluated retrospectively on AE datasets from three prospective AML clinical trials (MOSAIC, DELTA, and DaunoDouble) with automated LLT/PT/SOC metrics and expert-assessed Clinical Correctness Rate (CCR). Results: Clinical expert review showed high clinical acceptability of the RAG pipeline across datasets (91-97%). Under automated evaluation, the pipeline achieved LLT exact accuracy of 50-58%, PT accuracy of 78-85%, and SOC accuracy of 90-93%. Zero-shot generation and random candidate selection performed substantially worse. Semantic retrieval more often included the coder-assigned LLT among the candidate terms available to the model than retrieval based on lexical similarity. Multi-model benchmarking showed that backbone choice mainly affected LLT exact agreement, whereas PT and SOC performance remained comparatively stable. Conclusions: Retrieval-augmented reasoning supports clinically aligned MedDRA coding of free-text adverse events under realistic candidate constraints in AML clinical trials. Evaluation across three AML clinical trials showed that strict LLT-level string agreement underestimated clinical ap-propriateness, highlighting the importance of combining hierarchical evaluation metrics with clini-cal expert validation for AI-assisted MedDRA coding in hematology trials.

8
A Drug-Specific, Half-Life-Adjusted Framework for Classifying CNS-Active Systemic Therapy Exposure During and After Radiotherapy

Pari Mitre, L.; Drapkin, B.; Dohopolski, M.

2026-06-22 health informatics 10.64898/2026.06.11.26354463 medRxiv
Top 0.1%
13.0%
Show abstract

Clinical oncology datasets often store systemic therapy as a regimen label with a start date and an end date. Those records are clinically recognizable but can be analytically incomplete when the research question concerns whether a patient was exposed to a concurrent CNS-active drug (cCNS-aD) or an adjuvant CNS-active drug (aCNS-aD) around radiotherapy. Contemporary CNS-oncology studies usually define CNS activity by empiric drug lists and define concurrency by fixed calendar windows, although the literature shows substantial heterogeneity across both concepts. This paper proposes a generalizable framework for converting raw systemic therapy records into reproducible cCNS-aD and aCNS-aD variables, useful in subgrouping for clinical studies. The framework uses a transparent CNS scoring model based on three clinical evidence components: intracranial objective response rate, consensus CNS endorsement, and intrathecal route of administration. It then defines a pharmacokinetic exposure proxy as the recorded end date plus five half-lives. Concurrent exposure is classified by overlap with the radiotherapy interval, while post-radiotherapy exposure is classified by overlap with a prespecified post-RT attribution window. The framework separately identifies post-RT pharmacokinetic persistence and post-RT treatment initiation, allowing investigators to distinguish continued exposure from true adjuvant initiation. This is a methodological framework and reference implementation. Implementation audits and endpoint-specific sensitivity analyses remain necessary before use as a definitive exposure classifier

9
Adaptive Post-Processing Recovers Most of the Gap to nnU-Net v2 in Head and Neck GTV Segmentation: A Paired Three-Arm HECKTOR 2025 Benchmark

Oyarzun Silva, R.; Hernandez Hernandez, P.

2026-08-31 radiology and imaging 10.64898/2026.08.28.26361649 medRxiv
Top 0.1%
11.7%
Show abstract

Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.

10
Gate-Before-Generate: A Dual-Layer Architecture for Output-Presence Routing in Chest X-ray Report Generation

BAI, T.-C.; YEH, S.-C.

2026-08-12 radiology and imaging 10.64898/2026.08.11.26360223 medRxiv
Top 0.1%
10.0%
Show abstract

CXR report generation may require a vision-language model (VLM) to produce both textual findings and spatial bounding boxes. Generative 4B-7B VLMs can emit non-empty outputs on normal images and empty outputs on abnormal images, motivating explicit structural routing. To evaluate whether a hard inference-time gate before a probabilistic VLM changes output-presence performance and to identify the mechanisms underlying paired STRUCT outcomes. We evaluated CXRxVLM v2, combining a frozen microsoft/rad-dino ViT-B/14 encoder with a 768[-&gt;]1 logistic probe (threshold 0.0557) and google/medgemma-4b-it with the pamessina/medgemma-4b-it-cure LoRA adapter. A seed=42 stratified cohort of 500 VinDr-CXR train-pool images (250 NORMAL, 250 ABNORMAL) was compared with Lingshu-7B A_baseline and D_fewshot configurations. Exact paired McNemar tests and stratum-level output-presence analyses were prespecified for the primary configurations; MedGemma 1.5 SigLIP was exploratory. CURE achieved STRUCT = 78.0% (390/500; Wilson 95% CI 74.2-81.4), versus 73.8% for Lingshu A_baseline and 74.2% for D_fewshot. Pairwise p-values were 0.0778, 0.1042, and 0.8642. The paired decomposition showed CURE ABNORMAL non-empty-output advantage of +13.6 percentage points versus Lingshu A (p = 0.0012; +14.0 points versus D, p = 0.0007), while Lingshu had higher NORMAL empty-output rates (+5.2 to +6.4 points; p = 0.0106 and p = 0.0004). The full pipeline used 8.87 GB VRAM and 4.92 s/image mean latency; 53% of records used a 25.7 ms warm gate-negative path after model loading. Equivalent overall STRUCT scores concealed two mechanistically different output regimes: CURE favored ABNORMAL non-empty outputs, whereas Lingshu favored NORMAL empty outputs. This paired decomposition, rather than the aggregate score alone, characterizes how hard-gated and probabilistic systems route output presence.

11
Large language models for cancer registry abstraction: a real-world evaluation across models, variables, and cancer types

Fuchs, J.; Satusky, M. J.; Leese, P. J.; Nag, S.; Zipple, I. W.; Baggett, C. D.; Lash, S.; Reeder-Hayes, K.; Wood, W. A.; Johnson, C. T.; Critchley, C.; Krishnamurthy, A. K.; Elston Lafata, J.; Thompson, C. A.; Troester, M. A.; Pfaff, E. R.

2026-06-29 health informatics 10.64898/2026.06.25.26356626 medRxiv
Top 0.1%
9.6%
Show abstract

Cancer registries enable cancer surveillance at the population level. These registries require significant human-time to read through many different parts of the electronic health record, including structured data and lengthy, free-text clinical reports, to abstract values for hundreds of required variables. Large language models (LLMs) offer the possibility to significantly improve this process by supporting and speeding up cancer registry data abstraction. However, it is unclear how well these models perform at real-world cancer registry abstraction involving multiple cancer types and large patient volumes. Here, we evaluate five foundational LLMs for their ability to reliably abstract cancer registry variables. We leverage hospital cancer registry data from a large regional health system as the ground truth and use LLMs to abstract from clinical reports eight registry variables for 5,939 patients with seven different cancer types. We use a zero-shot prompting strategy to compare LLM ability on commonly abstracted cancer variables with different data types. The results show that larger and more advanced models (Claude Sonnet 4.5, GPT-OSS-120b, GPT-OSS-20b) generally outperform smaller models (Gemma 12b, LLaMA 3.1 8b). The best performing models show F1 scores around 0.8 for cancer registry variables with low cardinality (grade, summary stage, laterality), with only slightly lower F1 scores for variables with high cardinality (primary site, regional nodes examined, regional nodes positive). On the more complex task of precise date extraction, all models showed decreased performance on both diagnosis and treatment dates (exact accuracy ~0.55 for the best performing models), which increased to ~0.85 for a tolerance within {+/-}30 days. These results quantify the performance of various models as well as the potential and limitations of LLMs in cancer registry abstraction tasks.

12
Preoperative Prediction of Residual Cancer Burden After Neoadjuvant Chemotherapy in Breast Cancer: A Multimodal Machine Learning Approach and Implications for Clinical Decision Support

Dagdeviren, Y. K.; Semiz, H. S.; Inan, E. H.; Karakas, H. Y.; Durak, M. G.; Tezel, N.; Sevindik, M. C.; Kirmizibayrak, P. B.; Bekis, R.

2026-08-18 oncology 10.64898/2026.08.16.26360557 medRxiv
Top 0.1%
9.6%
Show abstract

Background. Residual cancer burden (RCB) after neoadjuvant chemotherapy (NAC) offers finer prognostic stratification than binary pathologic complete response, and increasingly guides adjuvant treatment intensity. Predicting four-tier RCB class from preoperative data could inform adjuvant planning before surgery, yet this remains an unmet need; and when two models reach equal discrimination, the key question is which generalizes most reliably. We compared a radiology-focused model with a fully integrated multimodal model for preoperative four-class RCB prediction. Methods. In a single-center, retrospective cohort of 328 patients treated with NAC followed by surgery, 64 clinicopathologic and radiologic variables were organized into thematic blocks. Two configurations were compared: a 17-variable radiology model (Model R) and a 62-variable multimodal model (Model ALL). Three algorithms (Random Forest, XGBoost, LightGBM) were evaluated with and without SMOTE using an 80/20 stratified split and 5-fold cross-validation. Model selection combined test AUC, macro-F1, cross-validation-to-test gap, nested cross-validation, bootstrap confidence intervals, and SHAP explainability, following the TRIPOD+AI guidance. Results. RCB classes were distributed as RCB-0 27.4% (n=90), RCB-I 10.4% (n=34), RCB-II 43.6% (n=143), and RCB-III 18.6% (n=61). Model R and Model ALL reached identical test AUC (0.838). Model ALL, however, achieved higher accuracy (0.636 vs 0.530) and macro-F1 (0.602 vs 0.598), together with a substantially smaller cross-validation-to-test gap (0.015 vs 0.099), pointing to more stable generalization; this gap difference persisted across all three algorithms. SHAP analysis showed that the multimodal model drew jointly on imaging phenotype, tumor biology, and disease extent. Both models remained weakest in the RCB-III class. Conclusions. At equivalent discrimination, the multimodal model was methodologically preferable for preoperative RCB prediction, owing to its stability and interpretability - qualities relevant to trustworthy clinical decision support. It remains investigational; a model flagging likely RCB-0 or RCB-III before surgery could prioritize adjuvant-therapy discussions earlier in the care pathway, pending prospective external validation.

13
megaMine: a scalable, rule-based framework for mining gene-cancer-drug evidence from biomedical literature

JUNAID, M.; Prazanowska, K. H.; Jeong, H.-E.; Ryu, Y.; Choi, J.; An, J.-Y.; Lim, S. B.

2026-08-12 bioinformatics 10.64898/2026.08.06.743392 medRxiv
Top 0.1%
8.0%
Show abstract

The rapid expansion of the oncology literature has outpaced manual curation of clinically relevant gene-cancer-drug associations and oncogenic driver evidence. Existing automated approaches often lack transparency or are difficult to scale across heterogeneous data sources. To address this gap, we developed megaMine, a transparent, rule-based, and context-aware literature-mining framework that integrates therapeutic and driver evidence from PubMed, PubTator, and Europe PMC by combining entity recognition, hierarchical heuristics, and contextual labeling. In therapy mode, megaMine was applied to approximately 100,000 oncology articles published between 2015 and 2025, yielding more than 23,000 structured sentence-level evidence records, with standardized annotations for drug response, resistance, and study context. Internal evaluation of context labels showed the strong separability between efficacy and non-efficacy evidence using ridge logistic regression (AUROC = 0.915; AUPRC = 0.941). Benchmarking against NCI/OncoKB-supported drug-cancer associations showed that curated clinical associations had higher megaMine composite evidence scores than unlabeled comparison pairs [median (IQR): 25.6 (9.07-72.5) vs. 3.61 (1.69-8.69); Wilcoxon rank-sum test, P < 2.2 x 10-16]. In driver mode, megaMine retrieved mutation- and biomarker-related evidence from an ERBB-focused gastric cancer query, generating 750 evidence rows from 200 PMIDs. These results demonstrate that deterministic and interpretable approaches can support scalable evidence extraction for downstream applications such as knowledge graph construction and literature-based evidence synthesis.

14
Variantscape: Large Language Model-Driven Mining of Biomedical Literature for Clinical Interpretation of Cancer Variants

Wosny, M.; Blindu, A. S.; Boesch, M.; Peres, T.; Niederhauser, T.; Fruh, M.; Rothermundt, C.; Hastings, J.

2026-08-03 health informatics 10.64898/2026.08.02.26359492 medRxiv
Top 0.1%
7.9%
Show abstract

Background: Precision oncology relies on accurate interpretation of tumour-detected gene variants, to guide personalized treatment decisions. However, accurate interpretation of variants in context requires extensive information that is often buried within unstructured biomedical literature and obscured by inconsistent nomenclature, making manual retrieval labour-intensive and prone to omissions. Methods: To address this challenge, we developed Variantscape, a large-scale, automated pipeline and open-access web tool. It integrates traditional natural language processing methods with state-of-the-art large language models to extract, standardize, and analyze co-associations between genetic variants, cancer types, and therapeutic interventions from published biomedical abstracts. Findings: From over 3 million abstracts screened, 335,817 gene name-containing articles were eligible for downstream extraction. Among these, 7,423 (2.2%) simultaneously mentioned a variant, cancer type, and therapeutic agent, encompassing 3,902 unique variants across 98 cancer types and 388 therapeutic agents. This highlights the inefficiency of manual literature retrieval in molecular tumour board (MTB) workflows. Network analysis revealed 14,831 statistically significant co-associations, represented in a literature-derived graph with 4,388 nodes and 46,943 edges. Canonical alterations in well-studied cancers (e.g., BRAF V600E in melanoma) were strongly linked to established treatments, while several rare variants also emerged with high-confidence literature support. Interpretation: By applying large language models to biomedical literature, Variantscape enables scalable, context-aware extraction of trilateral variant-treatment-cancer relationships. This approach supports early evidence synthesis/hypothesis generation, highlights underrecognized or rare associations, and offers a practical resource for accelerating discovery and supporting precision oncology research and translation. Unlike static databases, Variantscape is continuously updatable and leverages large language model-based inference to uncover putative associations without manual curation. Variantscape has the potential to support MTB workflows and translational research by rapidly revealing signals from underlying abstracts.

15
Multi-Timepoint Risk Stratification in Rare Cancers: A Computational Framework Validated against Published Ewing Sarcoma Trial Data

Kress, J.

2026-07-07 oncology 10.64898/2026.07.03.26357236 medRxiv
Top 0.1%
7.8%
Show abstract

Three audiences -- the family of a newly diagnosed Ewing sarcoma patient, the long-term survivor, and the cooperative-group trial statistician -- receive cohort-mean answers to patient-level questions because the patient-level data machine learning requires do not exist for rare cancers. We present a framework producing patient-level predictions from published aggregate trial data. A six-stage discrete-event Monte Carlo simulation integrates genetic risk factors, serial biomarker dynamics with genotype-conditional weighting, post-surgical ctDNA-based minimal residual disease (ctDNA-MRD) assessment, and treatment-related mortality as a separable competing risk. Adverse-effects modules project 30-year incidence across five organ systems from chemotherapy and radiation exposures. Its four structural ingredients are instantiated in Ewing sarcoma and validated against trial data from more than 3,400 patients. The framework achieves 3.2% mean absolute error across 23 efficacy endpoints (none exceeding 6%) and falls within published confidence intervals for all 20 toxicity endpoints. ctDNA-MRD stratification separates candidate populations -- 5.5% recurrence (de-escalation) versus 87.8% (intensification) -- and multi-timepoint integration produces 16-fold five-year EFS resolution spanning 5-96%, exceeding the 3- to 5-fold ranges of single-timepoint approaches. The 16.1-fold recurrence risk ratio emerges from simulation, not as a supplied parameter. Genotype-conditional weighting improves discrimination over equal-weight scoring in every subgroup (Pearson r +0.060 to +0.129), with largest gains where biological rationale is strongest. A Monte Carlo framework calibrated to published aggregate data turns cohort-mean answers into patient-level predictions as exemplified in the rare cancer Ewing sarcoma, where the conventional patient-level machine-learning pathway is structurally unavailable; transfer to other rare cancers remains a hypothesis for future validation. Survivorship-surveillance refinement is the most concrete current use; trial-design and prognostic counseling are next-decade pathways.

16
Deep Learning-Based Pretreatment Cardiovascular Risk Stratification in Women with Breast Cancer

Dehghan Manshadi, M.; Manouchehri, N.; Hubbert, L.; Liljegren, A.; Manouchehrinia, A.; Linder-Stragliotto, C.; Rantala, J.; Hedayati, E.; Kiani, N.

2026-07-28 epidemiology 10.64898/2026.07.26.26350307 medRxiv
Top 0.1%
7.4%
Show abstract

Background: Cardiovascular disease is a leading non-cancer cause of morbidity and mortality among breast cancer (BC) survivors. Existing cardiovascular risk tools are not tailored to cancer populations and often rely on cardiology investigations or treatment details unavailable at the initial oncology visit, limiting their use for early referral decisions. Methods: We conducted a registry-based cohort study including 17,051 women diagnosed with BC in stage I-III or ductal carcinoma in situ in the Stockholm-Gotland region (2008-2019). Using only pre-treatment information routinely available to oncologists, such as demographics, cancer characteristics, planned cancer treatment, baseline comorbidities, medications, and healthcare utilization, we trained and validated a deep learning-based competing-risk model to predict 1-year major adverse cardiovascular events (MACE) risk, accounting for non-cardiovascular death as a competing outcome. Model performance was evaluated using a 3-fold CV and time-dependent concordance indices. Fine-Gray subdistribution hazard models were used to aid interpretability. Results: The model demonstrated strong and stable discrimination across validation folds, with a median c-index of 0.84 for 1-year MACE prediction and 0.90 for the competing risk. Key contributors to predictive performance included age at BC diagnosis, prior cardiovascular disease, healthcare utilization patterns, cancer stage, and specific medication profiles. Several predictors with modest marginal hazard ratios in the Fine-Gray model contributed substantially through nonlinear effects and interactions. Conclusions: Our model with a deep learning-based competing-risk structure and using only pre-treatment, oncology-accessible data enables accurate short-term cardiovascular risk stratification in women with BC and may support targeted cardio-oncology referral prior to initiation of systemic therapy.

17
Calibrated Uncertainty Quantification for Patient-Level AML Drug Sensitivity Prediction Using Split Conformal Prediction

Shokrzadeh, A. J.; Shokrzadeh, P.

2026-06-11 bioinformatics 10.64898/2026.06.07.730728 medRxiv
Top 0.1%
6.7%
Show abstract

Accurate prediction of ex vivo drug sensitivity in acute myeloid leukemia (AML) patients from transcriptomic data is a critical challenge for precision oncology. Existing computational approaches have explored uncertainty quantification in cancer drug response prediction primarily using cell line data, while patient-level AML models typically rely on heuristic confidence measures rather than statistically calibrated uncertainty estimates. Here, we present a framework applying split conformal prediction to patient-level AML drug response modeling using the BeatAML 2.0 cohort. We trained Elastic Net and XGBoost regressors on bulk RNA-seq gene expression profiles from 318 AML patients, analyzing 34,764 patient-drug observations across 122 compounds. Baseline models achieved median Pearson R values of 0.291 (Elastic Net) and 0.281 (XGBoost) across 122 drugs. Wrapping these models with split conformal prediction yielded well-calibrated prediction intervals across three confidence levels: empirical coverages of 81.4%, 90.7%, and 95.5% against nominal targets of 80%, 90%, and 95%, respectively. Analysis of prediction interval widths revealed substantial drug-class-specific uncertainty patterns, with HDAC and BCL-2 inhibitors exhibiting markedly higher uncertainty than MDM2 inhibitors, suggesting a potential association between transcriptomic predictability and drug mechanism of action, although several drug classes were represented by only a small number of compounds. Predictive uncertainty was not significantly associated with ELN2017 molecular risk classification (Kruskal-Wallis p=0.395) or NPM1 mutation status (p=0.788). These results demonstrate that statistically valid uncertainty quantification can be achieved for patient-level AML drug response prediction despite substantial biological heterogeneity. to the best of our knowledge, no published study has applied split conformal prediction to patient-level ex vivo drug sensitivity prediction in the BeatAML cohort, providing a principled alternative to heuristic confidence scoring approaches.

18
IMMF: An Interpretable Multi-Modal Framework for Hypothesis-Driven Biomarker Discovery in Triple-Negative Breast Cancer Using Public Data

Imran, A.; Rahat Hossain, K. M.; Islam, S. M. R.; Rahman, M. S.

2026-08-24 bioinformatics 10.64898/2026.08.19.745809 medRxiv
Top 0.1%
6.6%
Show abstract

Triple-Negative Breast Cancer (TNBC) is characterized by high heterogeneity, poor prognosis, and limited targeted treatment options. Bridging the gap between molecular alterations and histopathological morphology remains a major challenge in precision oncology. We propose an interpretable, multi-modal framework that integrates histopathological image analysis with multi-omics profiling (somatic mutations, DNA methylation, copy number alterations), leveraging U-Net-based nuclei segmentation, vision-language models (BLIP), biomedical language models (BioGPT), and explainable AI (SHAP, LIME). Our framework achieves strong predictive performance (AUC = 0.989) and provides transparent, biologically grounded interpretations by integrating morphological features with genomically prioritized biomarkers. Cross-modal analysis confirms established TNBC drivers and generates novel, testable hypotheses associating specific epigenetic alterations with distinct morphological phenotypes. While causal validation requires future wet-lab experiments, our framework accelerates hypothesis-driven biomarker discovery by integrating complementary data modalities with language-based reasoning, providing a transparent foundation for hypothesis generation and clinical translation.

19
Pretrained transformers applied to population cancer registries improve survival prediction in label-scarce and previously unseen cancers

Gao, Y.; Yu, S.; Xia, Y.; Chen, S.; Xia, S.; An, R.; Zeng, J.; Zhao, F.; Ma, Y.; Wang, Y.; Xie, X.; Zhang, J.

2026-09-03 oncology 10.64898/2026.08.30.26361693 medRxiv
Top 0.1%
6.6%
Show abstract

Prognostic models in oncology are developed one cancer at a time, from that cancer's own labelled outcomes, and fail where prognostic information is scarcest. Rare cancers account for roughly a fifth of diagnoses and most paediatric malignancies, yet seldom supply enough events for a reliable time-to-event model. We therefore asked whether a representation learned without outcome labels can supply what those cohorts cannot. A Transformer encoder was pretrained by masked field-value modelling on 9425135 tumour records from the SEER 17 registries, diagnosed in 2000 to 2023. Only diagnosis-time fields passing a fail-closed coding-verification gate were admitted, and each record was emitted as an era-specific and a harmonised view, keeping two decades of recoding auditable. The encoder was then frozen and read by a linear Cox head for overall survival. Nine rare cancers were removed from the pretraining corpus entirely, each requiring an independent pretraining run. On a sealed test partition, all nine exceeded an architecture-identical random frozen encoder in Harrell concordance by +0.0034 to +0.0368, every lower confidence limit above zero. At 256 labelled patients, all 67 cancers favoured the pretrained representation over budget-matched Cox regression, median difference +0.0283. The advantage was bounded: given the entire training set, Cox regression was favoured in seven of nine rare cancers. The encoder did not outperform a field-frequency baseline on its own objective, so upstream reconstruction did not predict downstream transfer. Outcome-agnostic registry pretraining carries prognostic signal into cancers it has never seen, and is most useful where labels are fewest, without establishing clinical utility.

20
RadGuide AI: Development and Technical Evaluation of a General Nuclear Medicine Agent for Traceable Radiopharmaceutical Decision Support

Gu, X.; Zhu, H.; Zhong, F.; Teng, G.-J.

2026-07-10 radiology and imaging 10.64898/2026.07.09.26357614 medRxiv
Top 0.1%
6.4%
Show abstract

Background: Nuclear medicine and radiopharmaceutical development require coordinated radiochemistry, dosimetry, molecular imaging, radiation-safety and clinical decision processes. Current workflows remain fragmented, difficult to audit and poorly standardised for evaluating domain-specific AI support. Methods: We developed RadGuide AI, a nuclear medicine agent built around a traceable data-model-tool loop. Patent, literature and clinical-trial records were converted into 15,596 initial QA items; relevance screening, completeness checks, semantic deduplication and cross-validation retained 5,474 core QA items. MedGemma-27B-Instruct served as the foundation model and was adapted with LoRA. The system incorporated 55 MCP-wrapped tools covering radiopharmaceutical R&D, clinical decision support, imaging analysis and radiation-safety/dosimetry. Evaluation used a locked N=200 benchmark with predefined denominators, leakage control, expert scoring, statistical procedures, factuality audits and tool-execution metrics. Results: RadGuide-LLM achieved 88.5% answer accuracy (177/200; 95% CI, 83.3-92.2%) and a Macro-Average score of 21.5/25 (bootstrap 95% CI, 20.9-22.0), exceeding GPT-4o, DeepSeek-V3.2 and the base MedGemma model in this technical evaluation. Supplementary audits reported guideline compliance, terminology recall, knowledge coverage, tool-routing success and preclinical/phantom dosimetry agreement with explicit denominators and confidence intervals. Interpretation: RadGuide AI converts nuclear medicine queries into auditable retrieval, tool selection, calculation, verification and reporting workflows. The findings support technical feasibility, not definitive patient-level clinical validation; prospective multicentre studies and external benchmark release remain required before clinical deployment.